Journal of Speech, Language, and Hearing Research
● American Speech Language Hearing Association
Preprints posted in the last 90 days, ranked by how well they match Journal of Speech, Language, and Hearing Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.
Show abstract
This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.
Dorsi, J.; Sandberg, C.; Lacey, S.; Nygaard, L.; Sathian, K.
Show abstract
PurposeTo examine speech iconicity for shape in aphasia, we compared iconicity ratings from people with aphasia to those from neurologically intact individuals and evaluated how iconicity relates to phonological and semantic processing profiles in aphasia. MethodEleven people with aphasia and 11 age- and gender-matched neurologically intact participants rated how rounded or pointed 50 auditory pseudowords sounded using a 5-point scale. Ratings from participants with aphasia were compared to predicted iconicity ratings derived from reference ratings from prior work and to ratings from neurologically intact participants. For each participant with aphasia, correlations between individual ratings and predicted ratings were related to measures of phonological and semantic processing. ResultsRatings from people with aphasia were significantly correlated with both the predicted ratings and the ratings from neurologically intact participants. The strength of the correlation between individual ratings and predicted ratings did not differ significantly between groups, although there was a trend toward weaker correlations in the aphasia group. There were indications that greater language impairment was associated with greater disruption of iconicity ratings; in particular, deficits in phonological segmentation and semantic processing were associated with reduced sensitivity to shape iconicity. ConclusionThese findings suggest that sensitivity to shape iconicity is preserved in individuals with aphasia to varying degrees. The specific nature of language impairment appears to play an important role in determining iconicity processing in aphasia.
Lee, S. H.; Wang, S.; Varkanitsa, M.; Kiran, S.
Show abstract
Macrolinguistic discourse analysis offers valuable insight into how patients with neurogenic communication disorders organize and produce informative speech, yet it remains a largely manual and labor-intensive process. We report an automated pipeline for macrolinguistic discourse analysis for individuals with aphasia and dementia that integrates automatic speech recognition (ASR), utterance segmentation, sentence-level embeddings, centroid-based main-concept matching, and rule-based coherence error classification. These algorithms were applied to Cinderella story retellings from 309 participants (113 controls, 102 post-stroke aphasia (PWA), and 94 dementia). The algorithm reliably identified main concepts (83% accuracy against human labels) and derived interpretable features such as semantic distance to a main concept centroid, main concept coverage, and coherence error rates. Crucially, diagnostic classification results showed that logistic-regression classifiers trained on 10 macrolinguistic features distinguished aphasia from controls with high accuracy (AUC {approx} 0.94) but showed weaker separation for dementia (controls vs dementia AUC {approx} 0.66; aphasia vs dementia AUC {approx} 0.58). Semantic distance to the centroid emerged as a robust, informative predictor for diagnostic classification, demonstrating that the ability to produce narrative-aligned speech is clinically important. The automated pipeline enables scalable macrolinguistic discourse analysis that could support screening and longitudinal monitoring of discourse impairments across neurogenic populations.
Herrmann, B.; Fink, L. K.; Pandey, P. R.; Johnsrude, I.; Ryan, J. D.
Show abstract
Speech comprehension in noisy environments often requires cognitive effort, but listeners may disengage when comprehension becomes impossible. Eye movements have recently emerged as a promising new measure of listening effort, but it remains unclear whether eye movements are sensitive to the full effort profile across easy, difficult, and impossible speech comprehension. Across four experiments, participants listened to sentences at easy, difficult, and impossible levels of multi-talker background babble while pupil size and eye movements were recorded. Pupil size generally followed the expected inverted u-shaped effort profile: low for easy speech, maximal for difficult but still intelligible speech and lower again for impossible speech, although this pattern partly reflected sustained, condition-specific differences and not only sentence-evoked responses. Gaze dispersion - measuring the spread of eye movements - decreased with high temporal selectivity during difficult relative to easy and impossible speech, indicating reduced eye movements during active, effortful listening. However, gaze dispersion was also lower, but less temporally selective, during impossible compared to easy listening, especially in non-baseline-corrected analyses, suggesting that reduced eye movements do not index listening effort uniquely. Instead, eye movements appear to reflect both attentional engagement during difficult listening and disengagement or inward attention when meaningful listening is no longer possible. These findings indicate that pupil size and eye movements provide complementary indices of listening-related cognition, and highlight the integration of listening, cognition, and motor systems.
Colak, H.; Benzaquen, E.; Guo, X.; Lad, M.; Sedley, W.; Griffiths, T. D.
Show abstract
Understanding speech in noisy environments (SPIN) is an important everyday ability, and engaging in musical activities has been proposed as a factor that may support this ability. However, the cognitive mechanisms underlying a potential musical advantage in SPIN perception remain unclear. Here we investigated whether musical sophistication is associated with better SPIN perception in a large population-based sample, and whether this relationship is mediated by auditory working memory (AWM), verbal working memory (VWM), or non-verbal intelligence. We recruited 203 participants and measured SPIN perception at both word and sentence levels. Musical sophistication was assessed using the Goldsmiths Musical Sophistication Index (Gold-MSI). AWM was measured using delayed matching of tone frequency or the modulation rate of amplitude modulated white noise, VWM was based on backward digit span task, and non-verbal intelligence used matrix reasoning. Mediation analyses revealed that AWM fully mediated the relationship between musical sophistication and SPIN perception, whereas VWM showed no mediation effect. Non-verbal intelligence showed a partial mediating effect. Additional control analyses using structural equation modelling revealed that the indirect effect through AWM remained significant after accounting for age, hearing thresholds, and non-verbal intelligence. Together, these findings suggest that individuals with greater musical sophistication demonstrate better daily life listening abilities, and that superior auditory working memory may be the key cognitive mechanism underlying this advantage.
Wang, F.; Utianski, R. L.; Barnard, L. R.; Stricker, J. L.; Clark, H. M.; Meade, G. F.; Jones, D. T.; Whitwell, J. L.; Josephs, K. A.; Duffy, J. R.; Botha, H.
Show abstract
Motor speech disorders (MSDs) are early markers of neurological disease, but expert perceptual analysis is rarely available outside specialized centers. Automated speech analysis offers a scalable alternative, yet prior studies have not systematically compared modeling approaches or assessed clinically relevant metrics in independent datasets. This study compared static acoustic features, articulatory informed Phonet features, and self-supervised pretrained models for binary and multi label MSD classification. We trained and evaluated models on 583 speech samples using speaker level splits. Baseline models included logistic regression and Gated Recurrent Units (GRUs) trained on eGeMAPS and MFCCs. We extracted three types of Phonet derived features and evaluated pretrained HuBERT and SSAST models in frozen, partially fine-tuned, and fully fine-tuned configurations. Binary classification distinguished MSDs from controls, while multi label classification identified six MSD subtypes. Models were assessed using validation AUC, and cut points were tested on two independent datasets. Pretrained and Phonet based models substantially outperformed static acoustic features. In binary classification, HuBERT achieved the highest AUC (0.95), while compact Phonet derived GRUs achieved comparable performance (up to 0.94). These models generalized well to independent datasets, maintaining high sensitivity (0.94) and specificity (0.97). In multi label classification, Phonet models achieved the highest macro average AUC (0.86), but threshold-based subtype performance declined on unseen data. Automated MSD detection is feasible and clinically promising. Binary classification generalized well, whereas multi label classification showed limited threshold stability across datasets.
Guo, Z.-c.; McFarlane, K.; McHaney, J. R.; Choksi, I.; Feeney, M.; Preston, L.; Chandrasekaran, B.
Show abstract
ObjectivesObjective and ecologically valid measures of speech processing can complement conventional audiologic assessments. Phoneme-related potentials (PRPs), derived by averaging listeners electroencephalography (EEG) responses time-locked to phonemes in continuous speech, have emerged as a promising approach for capturing cortical processing of speech in naturalistic listening conditions. Importantly, PRPs reveal speech perception challenges even when conventional audiograms are clinically normal, positioning them as a promising neural marker for suprathreshold listening difficulties that standard audiometry often misses. As a critical step toward clinical translation, this study examined the extent to which PRP-derived measures remain stable across real-world contexts relevant to clinical implementation, including monaural versus binaural presentation, stimulus intensity level, and repeated testing sessions. The study also assessed cortical tracking of lower-level speech acoustics to determine whether the PRP findings could be attributed to acoustic processing. DesignEEG was recorded from 18 young adults with normal hearing as they listened to audiobook speech presented monaurally or binaurally at 60 or 75 dB across two sessions separated by approximately one week. Neural differentiation of phoneme manner-of-articulation classes (vowels, nasals/approximants, fricatives, and stops) in PRPs was quantified using two measures: an F-statistic reflecting between-manner relative to within-manner variability, and classification accuracy from a machine-learning model trained to predict manner class from PRPs. Temporal response function modeling assessed neural tracking of continuous acoustic envelope and onset features of the audiobook speech. ResultsNeither PRP-derived measure of manner differentiation showed significant effects of session, presentation modality, intensity level, or their interactions. Intraclass correlation analyses further indicated moderate-to-good reliability across all three factors. In contrast, neural tracking of the acoustic envelope and acoustic onsets was stronger under binaural than monaural presentation, with binaural presentation eliciting more pronounced cortical responses to the envelope. ConclusionsPRP-derived measures remained relatively stable across modest procedural variations that are common in clinical testing contexts, positioning PRPs as a potent objective index of naturalistic speech processing. This stability may reflect cortical processing of abstract, linguistically relevant speech categories and suggest that PRPs provide complementary information beyond audiologic assessments of peripheral auditory functions and EEG measures that primarily capture lower-level acoustic processing.
Marukatat, C.; Kaewrak, K.; Chunamchai, S.; Chunharas, C.
Show abstract
Spastic dysarthria diagnosis through subjective neurologist auditory-perceptual assessment remains standard practice despite known inaccuracy. To address this gap, we developed an objective framework grounded in phonetic evidence that spastic dysarthria preferentially impairs initial consonant articulation, using automatic speech recognition (ASR) to quantify dysarthria and localize corticobulbar lesions. We created four reading sentences targeting groups of initial consonants: labial (facial), lingual-alveolar (tongue), and velopharyngeal (pharyngeal/soft-palate) sentence, along with a mixed-consonant sentence for comparative evaluation. Thirty-seven patients with neuroimaging-confirmed corticobulbar lesions and 37 controls read each sentence. ASR transcribed dysarthric speech into text, and we computed a "syllable-error score" by counting incorrectly transcribed syllables. This yields a clinically meaningful feature that makes syllable-level phonetic errors explicit. Logistic regression models were trained for each sentence, and performance was summarized by the area under the receiver operating characteristic curve (AUC) across 10,000 resampled train-test splits. Consonant-specific sentences significantly outperformed the mixed sentence: the lingual-alveolar sentence performed best with (median AUC 0.88), followed by the labial (0.80), then the velopharyngeal sentence (0.72), while the mixed-consonant sentence was lowest (0.67). These results suggest that the interpretable ASR-derived syllable error feature, combined with a relevant machine learning classifier could inform clinical insight into consonant-specific vulnerability in spastic dysarthria, with lingual-alveolar consonants appearing particularly informative. Overall, this novel ASR-based framework, together with phonetics-informed feature design provides objective, accurate, and clinically meaningful digital quantification for spastic dysarthria detection and corticobulbar lesion localization.
Wade, N. E.; Bormann, B. M.; Mankel, K. M.; Comstock, D. C.; Das, S.; Whittle, R. S.; Brodie, H.; Sagiv, D.; Miller, L. M.
Show abstract
Pure tone audiometry (PTA) remains the clinical standard for evaluating hearing ability, yet individuals with similar audiometric profiles often exhibit substantial variability in their capacity to understand speech in everyday listening environments. Growing evidence suggests this variance is related to contributions from cognitive ability and auditory processing that standard threshold measures do not capture. To investigate how PTA, cognitive factors, and demographics such as age jointly predict real-world speech perception, 116 veteran adults 20-70 years old spanning a range of normal to moderate sensorineural hearing losses completed a spatial auditory attention task. Target color words were embedded within naturalistic short-story narratives presented under two conditions: a mono-talker speech-in-quiet (SIQ) condition and a dual-talker speech-in-noise (SIN) condition with a spatially separated competing narrative. Behavioral performance was quantified via color word hit accuracy, reaction time, and comprehension question accuracy. Participants also completed pure tone audiometry, the Montreal Cognitive Assessment (MoCA), and the Speech, Spatial and Qualities of Hearing Scale (SSQ12). Mixed-effects regression models were used to evaluate the contributions of PTA, age, cognitive ability, and self-reported hearing difficulty (SSQ12) to task performance across conditions. Results demonstrate a complex interplay between age, PTA, MoCA, and/or listening condition (SIQ vs. SIN) in predicting identification accuracy, reaction time, and comprehension. Age and condition significantly predicted hit accuracy and reaction time, with older participants showing improved accuracy in quiet but declining accuracy and slower responses in noise. PTA did not emerge as a significant main effect predictor but interacted with cognitive ability and condition to modulate performance, in some cases exhibiting a paradoxical inverse relationship with accuracy dependent on MoCA score. MoCA scores significantly predicted comprehension across conditions, and SIN hit accuracy was positively correlated with SSQ12 scores, validating the task against participants real-world listening experiences. These findings highlight the importance of incorporating cognitive screening and ecologically valid speech perception tasks into audiological assessment to better identify individuals at risk for functional hearing impairment in complex listening environments.
Lau, J. C. Y.; McHaney, J. R.; Goldman, L.; Robinshaw, K.; Mou, F.; McFarlane, K.; Chandrasekaran, B.; Losh, M.
Show abstract
Reported perceptual differences in autism may arise from reduced use of prior context to shape incoming sensory input. Speech perception provides a critical test of this account because stable perception requires listeners to integrate variable acoustic signals with contextual expectations. This study examined context-dependent modulation of speech encoding in autistic and non-autistic adults using the frequency-following response (FFR), a neurophysiological measure of phase-locked auditory encoding. Participants heard English intonational pitch contours presented in repetitive and variable contexts while EEG was recorded. Principal component analysis of FFR metrics yielded components indexing neural encoding fidelity and timing. Non-autistic participants showed enhanced encoding fidelity in more predictable contexts, whereas autistic participants showed reduced context-dependent modulation. Neural encoding timing also showed divergent context effects across groups, suggesting altered balance between feedback-based predictive mechanisms and locally driven adaptation processes. Within the autistic group, greater context-related modulation of encoding fidelity was associated with lower ADOS-2 Social Affect severity but poorer speech-in-noise perception, suggesting that the functional impact of contextual modulation depends on input reliability and task demands. These findings indicate that context-dependent modulation of speech encoding is altered in autism and may contribute to individual differences in auditory and social-communicative function.
Shao, M.; McNair, K. A.; Parra, G.; Tam, C.; Sullivan, N.; Senturk, D.; Gavornik, J. P.; Levin, A. R.
Show abstract
Individuals with autism spectrum disorder (ASD) often exhibit atypical auditory processing, yet it remains unclear whether and how the integration of simple acoustic features and contextual information is impacted in ASD. One real-world example of this integration is the auditory looming bias, the prioritized processing and perception of approaching auditory stimuli. We designed a paradigm that presents intensity-rising (looming) and intensity-falling (receding) auditory stimuli to 3-4-year-old children with ASD (n = 21), children with sensory processing concerns who do not have ASD (SPC; n = 16) and children with typical development (TD; n = 30). We recorded neural responses using electroencephalography (EEG) and found evidence of looming bias in the SPC and TD groups, as indexed by greater P1 peak amplitude during the looming than receding stimuli (TD: t(64) = 6.87, p < .001; SPC: t(64) = 4.07, p < .001). But this finding was not present in the ASD group (p = .194). Additionally, the ASD group showed reduced differentiation between looming and receding stimuli, as indicated by significantly lower Rise-Fall Difference Score (RFDS) in comparison to the TD group (Z = -3.00, padj = .008). These findings suggested altered context-dependent modulation of sensory input in ASD. Lay SummaryMany children with autism show differences in how they process sounds. Using sound patterns in which loudness gradually increased and decreased over time, like many real-world sounds, we found that children with autism showed less neural differentiation between increasing and decreasing sounds. This finding suggested that the brain may process changes in sound differently in autism, particularly in how it adjusts to sounds as they change over time, which could contribute to the sensory challenges many children with autism experience in daily life.
Galeano-Otalvaro, J.-D.; Dieudonne, B.; Francart, T.; Wouters, J.
Show abstract
Understanding speech in noisy environments relies strongly on binaural cues such as interaural time differences (ITDs) and interaural level differences (ILDs), which support spatial hearing and the segregation of competing sound sources. When these cues are degraded, listeners experience substantial difficulty in complex acoustic environments. Behavioural measures of binaural benefit, such as binaural masking level differences (BMLDs), binaural intelligibility level differences (BILDs), and spatial release from masking (SRM), are well established in normal-hearing (NH) listeners, but they require an active behavioural response. Neural speech tracking using electroencephalography (EEG) has emerged as a promising approach for quantifying neural processing of continuous speech, yet its sensitivity to spatial hearing cues remains insufficiently characterised. In this study, we investigated the neural correlates of spatial release from masking in NH listeners using EEG-based neural speech tracking. Nineteen participants listened to continuous Dutch speech stories presented with masking noise under two spatial configurations, collocated (S0N0) and spatially separated (S0N90), across multiple signal-to-noise ratios (SNRs). Neural tracking of the speech envelope was quantified using both envelope reconstruction and temporal response function (TRF) analyses. Spatial separation enhanced neural tracking of the target speech envelope, particularly at challenging SNRs where behavioural SRM was also observed. TRF analysis further revealed condition-dependent morphologies, including increased amplitudes and decreased latencies of late cortical components consistent with spatial unmasking effects. These neural differences were most pronounced at low SNRs, where spatial cues provide the greatest perceptual benefit. Together, these findings demonstrate that neural speech tracking captures cortical signatures of spatial unmasking and closely reflects behavioural improvements in speech understanding. Establishing these relationships in NH listeners supports the development of objective neural measures for evaluating binaural benefit in difficult-to-test populations.
Klis, A.;Menn, K.;Cetincelik, M.;Snijders, T.;Junge, C.
Show abstract
Speech consists of regularities at different timescales. Already during infancy, neural electrophysiological activity aligns to these rhythms. The degree to which infants exhibit neural tracking of speech can be linked to their language development. In this study, we examined how the neural tracking of sung speech develops across age, from infancy to early childhood, and across different frequency bands (i.e., at the stress, syllabic, and phonemic rates), and whether neural tracking at each frequency and age predicts childrens language outcomes. We included 2565 children of the longitudinal YOUth cohort. Children listened to Dutch sung nursery rhymes while EEG was recorded at three measurement waves. After preprocessing the data, we included 955 children at 5 months, 1048 children at 10 months, and 795 children at 2-4 years. The final sample consisted of 750 children who also completed a receptive vocabulary test at 2-4 years. Children from 5 months onwards showed significant neural tracking of stressed syllables, syllables, and phonemes, measured with speech-brain coherence (SBC). Unexpectedly, there were no developmental changes in SBC across different frequency bands from infancy to early childhood. As expected, children with larger receptive vocabularies showed increased SBC in the stressed syllable rate. These findings suggest that stronger tracking of stressed syllables is related to individual differences in language ability.
Sekine, K.; Okuma, R.; Ban, H.
Show abstract
People frequently gesture while speaking, even when listeners cannot see them--for instance, during phone calls or behind barriers. Congenitally blind individuals also gesture, indicating that gestures serve functions beyond visual communication. Previous models of gesture production (e.g., Kita & Ozyurek, 2003; Rauscher et al., 1996) suggest that gestures facilitate speech, but they rely heavily on behavioural data and provide limited insight into temporal dynamics. This study used magnetoencephalography (MEG), a neuroimaging technique with high temporal resolution, to investigate when gestures influence speech. Twenty-three native Japanese speakers took part in a storytelling task under two conditions: Gesture-Required (gesture use instructed) and Gesture-Prohibited (hands kept still). Participants described cartoon clips across multiple sessions (30 trials x 3 sessions per condition). Using speech onset as the reference point, we compared root mean square (RMS) values within a -0.25 to 0 second window. RMS values were higher in the Gesture-Prohibited condition, with increased activity in the bilateral anterior temporal lobes (Left ATL: p = .049; Right ATL: p = .027), but not in motor regions (p = .29). These findings suggest that gestures reduce neural load in language-related regions before articulation. Co-speech gestures may support speech planning by facilitating lexical retrieval or semantic structuring. The lack of motor region effects indicates that this influence is linguistic rather than motoric. This study provides direct direct neurophysiological evidence of the timing of gesture-speech interaction, supporting models that view gestures as an integral part of speech production.
Ip, E. Y. J.; Akkaya, A.; Winchester, M. M.; Bishop, S. J.; Cowan, B. R.; Di Liberto, G. M.
Show abstract
Human speech is inherently social. Yet our understanding of the neural substrates underlying continuous speech perception relies largely on neural responses to monologues, leaving substantial uncertainty about how social interactions shape the neural encoding of speech. Here, we bridge this gap by studying how EEG responses to speech change when the input includes a social element. In Experiment 1, we compared the neural encoding of synthesised undirected monologues, directed monologues, and dialogues. In Experiment 2, we extended this by using podcasts, addressing the additional challenges of real speech dialogue, such as dysfluency. Using temporal response function analyses, we show that the presence of a social component strengthens the cortical tracking of the speech envelope, despite identical acoustic properties. Neural responses to synthesised speech showed a strong correlation with those for real speech podcasts, with a stronger alignment emerging for more socially-relevant speech material. In addition, we demonstrate that robust neural indices of sound and lexical-level processing can be derived using real podcast recordings despite the presence of dysfluencies. Finally, we present a simulation to put to the test the robustness of temporal response function analyses under increasing levels of dysfluency. Together, these findings highlighting the impact of social elements in shaping auditory neural processing, providing a framework for future investigation and analysis of social speech listening and speech interaction. Significance StatementHuman speech is rarely produced or processed in a social vacuum. Yet, our understanding of continuous speech neurophysiology mostly comes from experiments involving speech monologues. This study reveals how social context modulates the neural encoding of speech. We directly contrast neural signals recorded when participants listened to monologues and dialogues, using controlled material from speech synthesis and real podcast recordings. We found that the social element amplifies the neural encoding of speech features, reflecting greater engagement. We also show strong correlation between synthetic and real podcast neural responses, scaling with social relevance. Finally, we demonstrate that lexical processing can be measured robustly even amid natural dysfluencies. These insights advance our understanding of speech neurophysiology, informing future research on social speech.
Prawiroharjo, P.; Fakhri, A.; Gabrielle, A.; Martalia, V.; Rahmayani, S. A.; Wijaya, V. G.
Show abstract
Aphasia diagnosis in Indonesia remains challenging due to limited culturally and linguistically appropriate instruments. Widely used tools such as the Boston Diagnostic Aphasia Examination (BDAE) and Western Aphasia Battery (WAB) are not adapted to the Indonesian context, while Tes Afasia untuk Diagnosis, Informasi, dan Rehabilitasi (TADIR) provides screening but lacks diagnostic accuracy. To address this gap, we developed the Instrumen Diagnosis dan Evaluasi Afasia (IDEA) for native Indonesian speakers and evaluated its validity, reliability, and normative cutoff values in cognitively healthy Indonesian adults. Eighty-three cognitively normal adults (screened using MoCA-Ina) with no history of neurological disease were assessed using IDEA, which evaluates six language domains. Items were adapted from existing tools and reviewed by experts. Content validity, internal consistency (Cronbachs alpha), and construct validity (Exploratory Factor Analysis) were analyzed using SPSS v25. A total of 83 participants were included (median age = 55.81 years, 54% secondary education). IDEA demonstrated good feasibility, with an average completion time of 45-60 minutes depending on participant engagement. Content validity was established by unanimous expert consensus. Construct validity showed meritorious sampling adequacy (KMO = .872) and significant sphericity (Bartletts test {chi}^2 (15) = 278.523, p<.001), supporting factor analysis. Internal consistency showed good reliability across six domains (Cronbachs = 0.896). IDEA is a valid and reliable tool for assessing aphasia in Indonesian natives. It is a culturally appropriate assessment tool which offers structured, domain-based evaluation and supports differential diagnosis of both classical and progressive aphasia syndromes. Keywords: Aphasia, Language Assessment, Indonesian, IDEA, Validity
Colak, H.; Guo, X.; Benzaquen, E.; Gurusiddappa, M.; Banerjee, A.; Choi, I.; Sedley, W.; Griffiths, T. D.
Show abstract
ObjectivesOutcomes following cochlear implantation vary substantially across adult recipients, and the cognitive and perceptual factors contributing to this variability are not fully understood. This poses a challenge for developing strategies to improve cochlear implant outcomes, as such approaches require a clearer understanding of the mechanisms underlying individual listening difficulties. In this study, we investigated auditory cognitive measures in cochlear implant (CI) users to further elucidate the origins of this variability. DesignThirty-seven adult cochlear implant users completed measures of auditory cognition, comprising auditory working memory (AWM) and sound segregation ability, measured using an auditory figure-ground task (AFG), as well as measures of peripheral temporal and spectral processing, comprising the temporal modulation detection threshold (TMDT) and spectral ripple discrimination threshold (SRDT). Speech perception outcomes were assessed using word-in-noise (WIN) and sentence-in-noise (SIN) tasks. Separate multiple linear regression models evaluated the unique contribution of the auditory cognition measures to WIN and SIN performance, after accounting for the peripheral measures. ResultsBoth regression models explained a substantial proportion of variance in speech-in-noise outcomes (WIN: adjusted R{superscript 2} = 0.55; SIN: adjusted R{superscript 2}=0.57, both p < 0.001). For WIN performance, AFG and AWM were significant predictors. A similar pattern was found for SIN performance, where lower AWM ability and poorer AFG segregation were linked to poorer sentence listening in noise. No significant effects of spectral ripple discrimination or temporal modulation detection were observed in either model, even though both were significantly correlated with WIN performance. ConclusionsThese findings indicate that auditory working memory and sound segregation ability are robust predictors of speech-in-noise outcomes in adult cochlear implant users, across both word- and sentence-level measures. Together, the results may help explain why speech-in-noise outcomes remain highly variable among CI users, even when basic sensory encoding abilities are taken into account. Incorporating measures of auditory working memory and fundamental sound segregation may therefore improve outcome prediction and help in developing more individualised rehabilitation strategies.
Wang, Z.; Li, G.; Yu, Y.; Wu, J.; Yu, Z.; Meng, Y.; Wang, S.; Dong, C.
Show abstract
Efficient face-to-face communication relies on the integration of auditory speech and visual articulatory signals. Over the past five decades, the McGurk illusion has been widely used as an index of audiovisual speech integration. However, substantial variabilities in susceptibility to the illusion across participants and speakers limit its reliability as a stable measure of audiovisual integration ability. Here, we introduce the McGurk illusion dataset (MID), which, to our knowledge, is the largest publicly available McGurk stimulus dataset to date. The MID comprises auditory (N = 400), visual (N = 400), and audiovisual (N = 640) speech stimuli generated from 80 Mandarin speakers and validated through behavioral judgments across 360,900 trials. Using this dataset, we characterized the acoustic and facial articulatory properties of McGurk stimuli, replicated substantial inter-participant and inter-speaker variabilities in illusion susceptibility, and revealed the associations between variations in McGurk illusion rate and the variations in unisensory perception, audiovisual correspondences, and speakers characteristics. Furthermore, the stimulus set enabled systematic comparisons of the reliability of different McGurk illusion-based indices of audiovisual speech integration. Overall, the MID not only provides a standardized resource for investigating audiovisual speech integration and its alterations across populations, but also supports research on speaker normalization, lip-reading, and speech perception.
Kwon, J.; Kotani, H.
Show abstract
During social interactions, people continuously align their movements and rhythms, a process known as interpersonal synchrony that supports rapport, mutual understanding, and smooth communication. In autism spectrum disorder (ASD), previous studies have often reported atypical or reduced synchrony, but most have relied on aggregate or session-averaged measures that may miss how coordination develops over time. It therefore remains unclear whether interactional differences in autism reflect a general reduction in synchrony or altered temporal dynamics of interpersonal coordination. We examined the temporal dynamics of head-movement synchrony during a structured face-to-face communication task, comparing non-autistic dyads (two typically developing [TD] partners) with mixed-neurotype dyads (one TD speaker paired with one autistic listener), using gyroscope-based tracking and time-resolved trajectory modelling. Phase-based synchrony, indexed by the phase-locking value (PLV), was lower overall in mixed-neurotype dyads. Critically, time-resolved analyses revealed a marked group difference in synchrony trajectories: non-autistic dyads showed progressive, adaptive growth in synchrony over the interaction, whereas mixed-neurotype dyads showed a significantly attenuated, flatter pattern. These findings suggest that autism may involve altered temporal organization of social coordination rather than simply reduced synchrony overall. Lay AbstractWhen we talk with someone, we often naturally match their body language and rhythms without even realizing it. This physical "syncing up" helps us feel connected, builds trust and shared understanding, and makes communication flow easily. Research shows that autistic people might sync their movements differently during conversations compared to non-autistic people. However, past studies usually just measured an overall average of this syncing across a whole interaction. This approach misses how human interactions actually unfold over time. We wanted to know: do autistic people just sync less overall, or does their syncing change differently as the conversation goes on? To find out, we used small motion sensors to track the head movements of adults having structured face-to-face conversations and compared two types of pairs: non-autistic pairs, where both people were non-autistic, and mixed-neurotype pairs, where one non-autistic speaker talked to one autistic listener. We found a notable difference in how the two groups interacted over time. For the non-autistic pairs, the physical syncing grew progressively stronger as the conversation progressed; they progressively "tuned in" to each other. In contrast, mixed-neurotype pairs showed a flatter pattern--their level of syncing stayed relatively constant from start to finish without that same gradual build-up. These findings are important because they suggest that differences in autistic communication are not simply a "lack" or "deficit" in social coordination. Instead, autistic individuals have a distinct style of interacting--one that maintains social engagement without relying on the progressive build-up of physical syncing that non-autistic people use. Taken together, our results highlight the importance of examining how interactions evolve over time to better understand the different ways autistic and non-autistic people communicate.
Liu, J.; Loudermilk, K.; Kim, K. S.
Show abstract
It has been demonstrated that people who stutter exhibit atypical motor control not only in speech tasks but also movements in the non-speech effector system, such as finger or arm motion. Notably, studies have reported that people who stutter show limited sensorimotor adaptation (i.e., updating subsequent movements in response to sensory errors) in both speech auditory-motor (i.e., updating speech movements in response to altered auditory feedback) and upper limb visuo-motor (i.e., updating arm movements in response to altered visual feedback) tasks. Given that speech auditory-motor adaptation is mostly if not entirely implicit (i.e., participants are unaware of the learning), it is thought that people who stutter have limited implicit adaptation in the speech effector system. It remains unclear however, whether such limited implicit learning also extends to upper limb visuomotor adaptation. Here, we examined implicit visuomotor learning in adults who stutter through the means of arm reaching adaptation to clamped visual feedback which provides a cursor that is fixed in direction (8{degrees} counterclockwise from targets) regardless of the participants actual hand location. All participants gradually adjusted their reach angle towards the clockwise direction, adapting in response to clamped feedback, but adults who stutter showed less adaptation compared to adults who do not stutter. In addition, computational modeling suggests that this implicit adaptation difficulties in stuttering individuals may reflect reduced error sensitivity. Together, our findings suggest that implicit sensorimotor learning difficulties in adults who stutter may generalize across multiple effector systems, providing important implications for understanding sensorimotor mechanisms underlying stuttering. Significance statementBy employing the clamped visual feedback paradigm during arm reaching movements, we demonstrated that adults who stutter showed less implicit visuomotor adaptation compared to adults who do not stutter. This study provides the first evidence that implicit sensorimotor adaptation limitations in developmental stuttering generalize across multiple effector systems. Our findings not only add to a growing body of evidence that stuttering is associated with domain-general sensorimotor difficulties but also point to specific underlying processes that may lead to stuttering.